Skip to content

Add complete documentation automation system - #3

Open
Cas-Dec wants to merge 1 commit into
h-cel:mainfrom
Cas-Dec:automation-system
Open

Cas-Dec wants to merge 1 commit into
h-cel:mainfrom
Cas-Dec:automation-system

Conversation

@Cas-Dec

@Cas-Dec Cas-Dec commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

This commit implements a comprehensive automation solution for HPC data documentation as discussed in the reform proposal.

New Features

Automation System

  • scripts/update_site.py: Main orchestrator that validates metadata, detects dataset changes (added/moved/deleted), rebuilds catalog, and commits changes
  • scripts/setup_cron.sh: Easy cron job setup and management
  • scripts/watch_and_update.py: Real-time file watcher alternative
  • State tracking via .automation-state.json for change detection
  • GitHub Actions workflow for CI/CD validation

Documentation

  • Reorganized all .md files into docs/ directory (except README.md)
  • Added QUICK_REFERENCE.md: one-page cheat sheet
  • Added STRUCTURE_DIAGRAM.md: visual guide with before/after
  • Added CONTRIBUTING.md: step-by-step contributor guide
  • Added QUICKSTART_AUTOMATION.md: 5-minute setup guide
  • Updated README.md: comprehensive, role-based entry point

Templates & Schema

  • schema/metadata-template.yaml: ready-to-use template
  • requirements-automation.txt: consolidated dependencies

Key Capabilities

✅ Automated metadata validation using jsonschema
✅ Dataset change detection (added/moved/deleted)
✅ Automatic catalog regeneration
✅ Cron-based or real-time file watching
✅ Production-ready with logging and error handling ✅ Dry-run mode for safe testing

🤖 Generated with Claude Code

@Cas-Dec
Cas-Dec force-pushed the automation-system branch 6 times, most recently from dd89815 to 5910d0e Compare October 6, 2026 14:21
This commit implements a comprehensive automation solution for HPC
data documentation as discussed in the reform proposal.

## New Features

### Automation System
- scripts/update_site.py: Main orchestrator that validates metadata,
  detects dataset changes (added/moved/deleted), rebuilds catalog,
  and commits changes
- scripts/setup_cron.sh: Easy cron job setup and management
- scripts/watch_and_update.py: Real-time file watcher alternative
- State tracking via .automation-state.json for change detection
- GitHub Actions workflow for CI/CD validation

### Documentation
- Reorganized all .md files into docs/ directory (except README.md)
- Added QUICK_REFERENCE.md: one-page cheat sheet
- Added STRUCTURE_DIAGRAM.md: visual guide with before/after
- Added CONTRIBUTING.md: step-by-step contributor guide
- Added QUICKSTART_AUTOMATION.md: 5-minute setup guide
- Updated README.md: comprehensive, role-based entry point

### Templates & Schema
- schema/metadata-template.yaml: ready-to-use template
- requirements-automation.txt: consolidated dependencies

## Key Capabilities

✅ Automated metadata validation using jsonschema
✅ Dataset change detection (added/moved/deleted)
✅ Automatic catalog regeneration
✅ Cron-based or real-time file watching
✅ Production-ready with logging and error handling
✅ Dry-run mode for safe testing

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <[email protected]>
@Cas-Dec
Cas-Dec force-pushed the automation-system branch from 5910d0e to c724ffa Compare October 6, 2026 14:34
@olivierbonte
olivierbonte self-requested a review October 7, 2026 08:23

@olivierbonte olivierbonte left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for giving this a first go @Cas-Dec !

Additional to the review comments, some more general comments:

  • the rendered site (now under docs-site/site) should not be added in version control. Github pages renders it from the .md files in the CI
  • The general feel of some parts of the website is a bit "claudish", as in a lot of text, some repetition... I would try to clean this up a bit. Therefore, I haven't yet reviewed the content under docs/, as a first glance not all seems relevant for the documentation website. The relevant parts can then be moved in to what is actually on the documentation website. Other parts can be removed, or added to a location that makes clear that these are only temporary notes.

Proposed structure:

|-- `docs` # merging of current `docs` and `docs-site`
|-- `data_catalog`
|     |-- VSC_DATA_VO # duplication of the data structure on VO, but only the .yaml files
|     |-- XXX # Potential future location where to collect data from
|-- `schema`
|-- `src` # Add here all code that is used for validating metadata, building catalog... as a package            

Bundling the code in src as package also has the advantage that you can easily make CLI tools out of it. @joppemassant does this very often in uv for GLEAM work, so you can always ask him for some advice on this.

runs-on: ubuntu-latest

steps:
- uses: actions/checkout@v3

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

update to newer version, currently already a v7: https://github.com/actions/checkout

Comment on lines +17 to +23
uses: actions/setup-python@v4
with:
python-version: '3.9'

- name: Install dependencies
run: |
pip install pyyaml jsonschema

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it is better to use 1 package manager across the project. Personal preference would be uv or pixi. uv and pixi can then be used both in CI and on HPC. I think that pixi is more HPC stable, but to be discussed.
Then you don't have to manually install with pip in CI

Comment thread README.md
# test_documentation
Repo to test MkDocs for internal documentation

<<<<<<< HEAD

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Merge conflict here

Comment thread README.md
- **Automated documentation** that stays in sync with the actual data
- **Clear folder structure** separating external data, processed data, projects, and personal files

## Quick Links

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think the separation bewteen docs and docs-site should stay. All info should be in 1 place, and all should be hosted on the website.

Comment thread build_and_serve.sh

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Instead of having a bash script for this, I would just document the 3 options:

  • mkdocs build
  • mkdocs serve
  • mkdocs build --strict
    or whatever their equivalents might be in the case of another static site tool

Comment thread scripts/update_site.py
contact_email = meta.get("contact", {}).get("email")
if contact_email:
logger.info(f" Would notify {contact_email} about {path.name}")
# TODO: Implement actual email sending via HPC mail system

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree, I think ideally we'd have a separate e.g. gmail account for this. This way, nobody has to store their personal password as a secret in the repo. https://docs.python.org/3/library/smtplib.html seems useful

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

as mentioned earlier: unless a clear advantage, I'd try to stick to one scheduling framework

@@ -0,0 +1,120 @@
{
"$schema": "http://json-schema.org/draft-07/schema#",
"$id": "https://h-cel.github.io/hpc-docs/schema/dataset-schema.json",

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

url not correct

"title": "Dataset metadata",
"description": "Schema for the metadata.yaml that must sit alongside every dataset folder under $VSC_DATA_VO/shared/.",
"type": "object",
"required": ["dataset", "source", "coverage", "variables", "contact", "history"],

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is nowhere in the docs this explicitly I think?

"type": "object",
"required": ["dataset", "source", "coverage", "variables", "contact", "history"],
"additionalProperties": false,
"properties": {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This properties field is not in the .yaml?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants